Papers with multimodal language models

14 papers
Compact Multimodal Language Models as Robust OCR Alternatives for Noisy Textual Clinical Reports (2026.eacl-industry)

Copied to clipboard

Challenge: Conventional OCR systems perform poorly under noisy, real-world conditions . compact multimodal models outperform classical and neural OCR pipelines .
Approach: They evaluate compact multimodal language models for transcribing noisy medical documents . they compare eight different models to find the best transcription accuracy and noise sensitivity .
Outcome: The proposed models outperform classical and neural OCR pipelines in transcription accuracy, noise sensitivity, numeric accuracy and computational efficiency.
The Impact of Auxiliary Patient Data on Automated Chest X-Ray Report Generation and How to Incorporate It (2025.acl-long)

Copied to clipboard

Challenge: Traditionally, CXR report generation relies on data from a patient’s exam, overlooking valuable information from patient electronic health records.
Approach: They propose to integrate patient data from ED records into multimodal language models that embed patient data into a language model.
Outcome: The proposed model incorporates patient data from the MIMIC-CXR and MIMICIV-ED datasets to improve diagnostic accuracy and improves radiologist effectiveness.
FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts (2024.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for visual question answering lack in visual grounding and complexity, particularly in evaluating spatial reasoning skills.
Approach: They propose to use flowcharts as visual contexts to assess the capabilities of visual question-answering multimodal language models in reasoning.
Outcome: The proposed benchmarks evaluate models' ability to follow visual information without pre-existing knowledge on a suite of open-source and proprietary multimodal language models using various strategies, followed by an analysis of directional bias.
A Computational Approach to Visual Metonymy (2026.eacl-long)

Copied to clipboard

Challenge: Visual metonymy is a form of indirect representation in which an image evokes a concept not by depicting it directly, but by presenting visually associated cues that invite the viewer to infer the intended meaning.
Approach: They propose a pipeline grounded in semiotic theory that leverages large language models and text-to-image models to generate metonymic visual representations.
Outcome: The proposed pipeline exploits large language models and text-to-image models to generate metonymic visual representations.
\mathsf{Con Instruction}: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities (2025.acl-long)

Copied to clipboard

Challenge: Existing attacks communicate instruction through text, accompanied by a toxic image or audio . a novel gray-box attack method generates adversarial images or audio to convey harmful instructions to MLLMs .
Approach: They propose a gray-box attack method that generates adversarial images or audio to convey specific harmful instructions to MLLMs by following non-textual instruction.
Outcome: The proposed method achieves highest success rates on visual and audio-language models . larger models are more susceptible toCon Instruction, compared to their underlying models - the results will be released .
Fine-Grained Prediction of Reading Comprehension from Eye Movements (2024.emnlp-main)

Copied to clipboard

Challenge: a new study attempts to assess reading comprehension from eye movements in reading . eye movements provide small improvements over a text-only baseline, the authors argue .
Approach: They propose to use eyetracking data to predict reading comprehension of a single participant . they use a battery of recent models and three new multimodal language models .
Outcome: The proposed model can predict reading comprehension of a single participant from eye movements over a paragraph.
Sharper and Faster mean Better: Towards More Efficient Vision-Language Model for Hour-scale Long Video Understanding (2025.acl-long)

Copied to clipboard

Challenge: Existing multimodal large language models (LLMs) have shown impressive performance on the video understanding task, but extremely long videos still pose significant challenges to their context length, memory consumption, and computational complexity.
Approach: They propose a vision-language model named Sophia for long video understanding which can efficiently handle hour-scale long videos.
Outcome: The proposed model exhibits competitive performance compared to existing video understanding baselines across various benchmarks for long video understanding with reduced time and memory consumption.
DeepInsert: Early Layer Bypass for Efficient and Performant Multimodal Understanding (2026.eacl-long)

Copied to clipboard

Challenge: Recent work shows that hyperscaling of data and parameter count in LLMs is yielding diminishing improvement when weighed against training costs.
Approach: They propose to insert multimodal tokens directly into the middle of the model to bypass the early layers.
Outcome: The proposed method reduces training and inference costs while preserving performance.
AlgoPuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Algorithmic Multimodal Puzzles (2025.naacl-long)

Copied to clipboard

Challenge: Existing datasets focused on visual question-answering focus on visual, language, and algorithmic knowledge . a new study examines the performance of multimodal language models in solving algorithmic puzzles .
Approach: They propose a dataset to test the capabilities of multimodal language models in solving algorithmic puzzles.
Outcome: The proposed dataset is generated automatically from human code.
Modeling Bottom-up Information Quality during Language Processing (2025.emnlp-main)

Copied to clipboard

Challenge: Contemporary theories of language processing model language processing as integrating both top-down expectations and bottom-up inputs.
Approach: They propose an information-theoretic operationalization for the “quality” of bottom-up information as the mutual information between visual information and word identity.
Outcome: The proposed model compares reading times in English and Chinese in which words' information quality has been reduced by occluding their top or bottom half with full words.
Language-Informed Synthesis of Rational Agent Models for Grounded Theory-of-Mind Reasoning On-the-fly (2025.findings-emnlp)

Copied to clipboard

Challenge: Language is a powerful source of information in social settings, especially in novel situations where language can provide both abstract information about the environment dynamics and concrete specifics about an agent that cannot be easily visually observed.
Approach: They propose a language-informed rational agent synthesis framework that integrates linguistic and visual inputs to draw context-specific social inferences.
Outcome: The proposed framework outperforms ablations and state-of-the-art models on a range of social reasoning tasks derived from cognitive science experiments.
Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models? (2025.findings-emnlp)

Copied to clipboard

Challenge: Despite significant progress in multimodal language models, it remains unclear whether visual grounding enhances their understanding of embodied knowledge compared to text-only models.
Approach: They propose to assess vision-language models’ perceptual abilities across different sensory modalities through vector comparison and question-answering tasks with over 1,700 questions.
Outcome: The proposed benchmark assesses the models’ perceptual abilities across different sensory modalities through vector comparison and question-answering tasks with over 1,700 questions.
Evo-PI: Aligning Medical Reasoning via Evolving Principle-Guided Supervision (2026.acl-long)

Copied to clipboard

Challenge: Existing models with static prompts, rules, or reward models are constrained by static supervision, which often fails to shape the underlying reasoning process, leading to brittle generalization and performance saturation in complex decision-making tasks.
Approach: They propose a principle-centric learning framework that treats reasoning principles as explicit, language-based supervision signals that can be generated, evaluated, and iteratively evolved.
Outcome: The proposed framework treats reasoning principles as explicit, language-based supervision signals that can be generated, evaluated, and iteratively evolved.
The Sound of Syntax: Finetuning and Comprehensive Evaluation of Language Models for Speech Pathology (2025.emnlp-main)

Copied to clipboard

Challenge: State-of-the-art multimodal language models (MLMs) show promise for supporting SLPs, but their use remains underexplored due to a limited understanding of their performance in high-stakes clinical settings.
Approach: They propose a taxonomy of real-world use cases of multimodal language models in speech-language pathologies to address this gap.
Outcome: The proposed model outperforms 15 state-of-the-art models in speech-language pathologies across five use cases and achieves improvements of over 30% on domain-specific data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations